Papers with Linguistic Data Consortium
Introducing NIEUW: Novel Incentives and Workflows for Eliciting Linguistic Data (L18-1)
Copied to clipboard
| Challenge: | a 2010 survey found that the language of the European Union, not even English, was not fully supplied . the absence of Language Resources stifles teaching and technology building, authors say . |
| Approach: | They propose to harness the power of alternative incentives to elicit linguistic data and annotation . they also describe changes to the workflows necessary to collect data from workforces attracted by incentives . |
| Outcome: | a new initiative to harness incentives to elicit linguistic data and annotation improves language resources . the NIEUW project is funded by the u.s. national science foundation . |
Reflections on 30 Years of Language Resource Development and Sharing (2022.lrec-1)
Copied to clipboard
| Challenge: | Linguistic Data Consortium was founded in 1992 to solve the problem that limitations in access to shareable data was impeding progress in Human Language Technology research and development. |
| Approach: | They review the roles of the Linguistic Data Consortium over the past 30 years after describing the conditions that lead to an HLT winter followed by a reawakening and an insatiable hunger for LRs. |
| Outcome: | The authors review the roles of the Linguistic Data Consortium over the past 30 years and provide a preview into future plans. |
CAMIO: A Corpus for OCR in Multiple Languages (2022.lrec-1)
Copied to clipboard
| Challenge: | CAMIO is a corpus of 70,000 images of machine printed text for optical character recognition (OCR) it covers 35 languages across 24 unique scripts. |
| Approach: | CAMIO is a corpus of annotated multilingual images for optical character recognition . the corpus includes nearly 70,000 images of machine printed text . |
| Outcome: | The corpus includes nearly 70,000 images of machine printed text . most images have been exhaustively annotated for text localization . |
Laying the Groundwork for Knowledge Base Population: Nine Years of Linguistic Resources for TAC KBP (L18-1)
Copied to clipboard
| Challenge: | Knowledge Base Population (KBP) evaluations target information extraction technologies for knowledge bases comprised of entities, relations, and events. |
| Approach: | They describe the linguistic resources provided by Linguistic Data Consortium for TAC KBP since 2009 . they highlight changes made to support evolving evaluation requirements . |
| Outcome: | The evaluations have targeted information extraction technologies for the population of knowledge bases comprised of entities, relations, and events. |
A 2nd Longitudinal Corpus for Children’s Writing with Enhanced Output for Specific Spelling Patterns (L18-1)
Copied to clipboard
| Challenge: | IQB study looks at reading, mathematics and spelling ability across different states. |
| Approach: | They collect three longitudinal corpora of German school children's weekly writing in German and transcribe them into a corpus for research via Linguistic Data Consortium. |
| Outcome: | The corpus of German school children's weekly writing in German was collected and transcribed. |
Related Works in the Linguistic Data Consortium Catalog (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing metadata standards for Related Works are used to define relations between language resources. |
| Approach: | They describe the development and implementation of a Related Works schema and the steps to implementation. |
| Outcome: | The proposed schema has been implemented in the Linguistic Data Consortium's (LDC) Catalog. |
A Progress Report on Activities at the Linguistic Data Consortium Benefitting the LREC Community (2020.lrec-1)
Copied to clipboard
Christopher Cieri, James Fiumara, Stephanie Strassel, Jonathan Wright, Denise DiPersio, Mark Liberman
| Challenge: | Linguistic Data Consortium (LDC) activities include the collection, annotation, processing, distribution, archiving and curation of language resources. |
| Approach: | a new report sketches the activities of a data center devoted to supporting the work of LREC attendees . 96 new corpora released in 2018-2020 to date, a technology evaluation campaign and innovations to advance methodology for language data collection and annotation. |
| Outcome: | 96 new corpora released in 2018-2020 to date, new technology evaluation campaign and innovations to advance methodology of language data collection and annotation. |
Morphological Segmentation for Low Resource Languages (2020.lrec-1)
Copied to clipboard
Justin Mott, Ann Bies, Stephanie Strassel, Jordan Kodner, Caitlin Richter, Hongzhi Xu, Mitchell Marcus
| Challenge: | a new corpus of annotated morphological data is described for the DARPA LORELEI Program . the data is annotating 9 low resource languages and root information for 7 of the languages . |
| Approach: | This paper describes a new morphology resource created by Linguistic Data Consortium and the University of Pennsylvania for the DARPA LORELEI Program. |
| Outcome: | The annotated corpus provides a gold standard for unsupervised morphological segmenters and analyzers . the language-specific annotation guidelines were language-independent, but included morphology paradigms and other specifications. |
From ‘Solved Problems’ to New Challenges: A Report on LDC Activities (L18-1)
Copied to clipboard
Christopher Cieri, Mark Liberman, Stephanie Strassel, Denise DiPersio, Jonathan Wright, Andrea Mazzucchi
| Challenge: | This paper reports on the activities of the Linguistic Data Consortium . |
| Approach: | This paper reports on the activities of the Linguistic Data Consortium . it summarizes the over 100 Language Resources released since the last report . |
| Outcome: | The report summarizes the over 100 Language Resources released since the last report . many of the LRs have been contributed by research groups around the world . |
The SAFE-T Corpus: A New Resource for Simulated Public Safety Communications (2020.lrec-1)
Copied to clipboard
| Challenge: | Linguistic Data Consortium developed the SAFE-T Corpus to support the NIST OpenSAT evaluation series. |
| Approach: | They introduce a new resource, the SAFE-T Corpus, designed to simulate first-responder communications by inducing high vocal effort and urgent speech with situational background noise. |
| Outcome: | The SAFE-T Corpus was developed to support the NIST OpenSAT (Speech Analytic Technologies) evaluation series. |
A Large Scale Speech Sentiment Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing corpus for sentiment analysis uses text inputs, but voice inputs are becoming more important as smart assistants and mobile voice control become more prevalent. |
| Approach: | They propose to extend the Switchboard-1 Telephone Speech Corpus by adding sentiment labels from 3 different human annotators for every transcript segment. |
| Outcome: | The proposed corpus contains 49500 labeled speech segments covering 140 hours of audio. |
Spanless Event Annotation for Corpus-Wide Complex Event Understanding (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing methods for annotating multilingual, multimedia data are limited by the availability of multilingual corpora for schema-based event representation. |
| Approach: | They propose a new approach to event annotation to promote whole-corpus understanding of complex events in multilingual, multimedia data. |
| Outcome: | The proposed method is part of the DARPA Knowledge-directed Artificial Intelligence Reasoning Over Schemas (KAIROS) Program. |